2 |
The CLASSLA-StanfordNLP model for lemmatisation of standard Slovenian 1.4
|
|
|
|
Abstract:
The model for lemmatisation of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the Sloleks inflectional lexicon (http://hdl.handle.net/11356/1230). The estimated F1 of the lemma annotations is ~99.7. The difference to the previous version of the model is that the Sloleks inflectional lexicon is moved to the morphosyntactic model.
|
|
Keyword:
language model; lemmatisation
|
|
URL: http://hdl.handle.net/11356/1478
|
|
BASE
|
|
Hide details
|
|
4 |
The Twitter user dataset for discriminating between Bosnian, Croatian, Montenegrin and Serbian Twitter-HBS 1.0
|
|
|
|
BASE
|
|
Show details
|
|
8 |
The news dataset for discriminating between Bosnian, Croatian and Serbian SETimes.HBS 1.0
|
|
|
|
BASE
|
|
Show details
|
|
9 |
The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.3
|
|
|
|
BASE
|
|
Show details
|
|
10 |
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild ...
|
|
|
|
BASE
|
|
Show details
|
|
13 |
Retweet communities reveal the main sources of hate speech
|
|
|
|
In: PLoS One (2022)
|
|
BASE
|
|
Show details
|
|
14 |
The ParlaMint corpora of parliamentary proceedings
|
|
|
|
In: Lang Resour Eval (2022)
|
|
BASE
|
|
Show details
|
|
18 |
Choice of plausible alternatives dataset in Croatian COPA-HR
|
|
|
|
BASE
|
|
Show details
|
|
19 |
Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0
|
|
|
|
BASE
|
|
Show details
|
|
20 |
The Orange workflow for observing collocation trends ColTrend 1.0
|
|
|
|
BASE
|
|
Show details
|
|
|
|